Demand MemCpy: Overlapping of Computation and Data Transfer for Heterogeneous Computing

نویسندگان

چکیده

Heterogeneous computing relies on collaboration among different types of processors shared data. In systems with discrete accelerators (e.g., GP-GPU), data sharing requires transferring a large amount between CPU and accelerator memories can significantly increase the end-to-end execution time. This paper proposes novel mechanism called Demand MemCpy (DMC) to hide overheads. DMC copies from host memory based demands at page granularity. It utilizes hardware-only fetch requested short latency background pre-copy related pages in advance. Our evaluation shows that reduce time GP-GPU application by 25.4% average overlapping computation transfer not unused pages.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Asymptotic algorithm for computing the sample variance of interval data

The problem of the sample variance computation for epistemic inter-val-valued data is, in general, NP-hard. Therefore, known efficient algorithms for computing variance require strong restrictions on admissible intervals like the no-subset property or heavy limitations on the number of possible intersections between intervals. A new asymptotic algorithm for computing the upper bound of the samp...

متن کامل

Automatic Transformation for Overlapping Communication and Computation

Message-passing is a predominant programming paradigm for distributed memory systems. RDMA networks like infiniBand and Myrinet reduce communication overhead by overlapping communication with computation. For the overlap to be more effective, we propose a source-tosource transformation scheme by automatically restructuring message-passing codes. The extensions to control-flow graph can accurate...

متن کامل

Accelerating Complex Data Transfer for Cluster Computing

The ability to move data quickly between the nodes of a distributed system is important for the performance of cluster computing frameworks, such as Hadoop and Spark. We show that in a cluster with modern networking technology data serialization is the main bottleneck and source of overhead in the transfer of rich data in systems based on high-level programming languages such as Java. We propos...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

ژورنال

عنوان ژورنال: IEEE Access

سال: 2022

ISSN: ['2169-3536']

DOI: https://doi.org/10.1109/access.2022.3195271